Papers by Pius von Däniken
TRANSLIT: A Large-scale Name Transliteration Resource (2020.lrec-1)
Copied to clipboard
| Challenge: | Transliteration is the process of expressing a proper name from a source language in the characters of a target language. |
| Approach: | They present a large-scale corpus of transliterated names in 180 languages . they use machine learning to train automatic transliteration . |
| Outcome: | The proposed system achieves 92% accuracy on identification of transliterated pairs. |
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing efforts in misinformation detection focus on written text, leaving a significant gap in addressing the complexity of spoken text in video transcripts. |
| Approach: | They propose to annotate video transcripts in three languages and six topics using a custom annotation tool. |
| Outcome: | The proposed tool shows strong cross-validation performance but challenges for generalization to unseen topics. |
On the Effectiveness of Automated Metrics for Text Generation Systems (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation methods lack a sound theoretical foundation for evaluation campaigns . imperfect automated metrics and insufficiently sized test sets are some of the factors that cause uncertainty. |
| Approach: | They propose a theoretical framework that incorporates different sources of uncertainty, such as imperfect automated metrics and insufficiently sized test sets. |
| Outcome: | The proposed model can be leveraged to improve evaluation protocols regarding reliability, robustness, and significance of the evaluation outcome. |
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems (2020.emnlp-main)
Copied to clipboard
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
| Challenge: | Lack of time efficient and reliable evalu-ation methods is hampering the development of conversational dialogue systems (chatbots). |
| Approach: | They propose a framework that replaces human-bot conversations with conversations between bots and an annotation tool that ranks chatbots based on their ability to mimic human behaviour. |
| Outcome: | The proposed evaluation framework replaces human-bot conversations with bot conversations and allows for frequent evaluations of chatbots during their evaluation cycle. |
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)
Copied to clipboard
| Challenge: | Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences. |
| Approach: | They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats. |
| Outcome: | The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral. |
Correction of Errors in Preference Ratings from Automated Metrics for Text Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation methods are over-confident in assigning significant differences between systems . Currently, the most reliable evaluation methods for text generation are human-based evaluations. |
| Approach: | They propose to combine human ratings with automated ratings to reduce the amount of human ratings needed to arrive at robust results. |
| Outcome: | The proposed evaluation protocol reduces the amount of human ratings by 50% while yielding the same evaluation outcome as the pure human evaluation in 95% of cases. |
LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts (2020.lrec-1)
Copied to clipboard
| Challenge: | Contractual provisions are a primary research target in law studies as they constitute the legal essence of a contract. |
| Approach: | They propose to use LEDGAR to construct a multilabel corpus of legal provisions in contracts that is crawled and scraped from the public domain. |
| Outcome: | The proposed corpus is the first freely available corpus of its kind. |
Do NOT Classify and Count: Hybrid Attribute Control Success Evaluation (2026.eacl-long)
Copied to clipboard
| Challenge: | evaluating attribute control success in controllable text generation relies on pretrained classifiers. |
| Approach: | They propose a Bayesian method that combines classifier predictions with a small number of human labels for calibration. |
| Outcome: | The proposed method produces robust estimates across both text and image generation tasks, offering an alternative to current evaluation practices. |